Papers with tri-modal architecture
Investigating Audio, Video, and Text Fusion Methods for End-to-End Automatic Personality Prediction (P18-2)
Copied to clipboard
| Challenge: | Using stacked Convolutional Neural Networks, we can predict personality traits from video clips with different channels for audio, text, and video data. |
| Approach: | They propose a tri-modal architecture to predict Big Five personality trait scores from video clips with different channels for audio, text, and video data. |
| Outcome: | The proposed model outperforms the best individual modality with 9.4% accuracy over the best channel. |